Papers with bilinear baselines
MIRTT: Learning Multimodal Interaction Representations from Trilinear Transformers for Visual Question Answering (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing bilinear methods focus on inter-modality information between images and questions . existing models focus on the interaction between images, questions, and images . |
| Approach: | They propose a trilinear interaction framework that incorporates attention mechanisms for capturing inter-modality and intra-modal relationships. |
| Outcome: | The proposed model outperforms bilinear models on the Visual7W Telling task and VQA-1.0 Multiple Choice task and outperformed baselines on the VQA, TDIUC and GQA datasets. |